Papers with language and vision tasks
Localizing Moments in Video with Temporal Language (D18-1)
Copied to clipboard
| Challenge: | a novel model for localizing moments in a longer video using natural language queries is challenging . moment localization is similar to other language and vision tasks, but it offers an interesting opportunity to model temporal dependencies and reasoning in text. |
| Approach: | They propose a model that explicitly reasons about different temporal segments in a video . their dataset includes a dataset with real videos and template sentences . |
| Outcome: | The proposed model explicitly reasons about different temporal segments in a video . it shows that temporal context is important for localizing phrases which include temporal language . |
Be Different to Be Better! A Benchmark to Leverage the Complementarity of Language and Vision (2020.findings-emnlp)
Copied to clipboard
| Challenge: | BD2BB is a language and vision benchmark that requires multimodal models combine complementary information from the two modalities. |
| Approach: | They propose a novel language and vision benchmark that requires multimodal models combine complementary information from both modalities. |
| Outcome: | The proposed model is easy for humans, but poor for humans . it compares state-of-the-art models against human speakers to show that it performs well. |
Cross-Modality Relevance for Reasoning on Language and Vision (2020.acl-main)
Copied to clipboard
| Challenge: | Existing approaches to learn and reason over language and vision data for downstream tasks such as visual question answering (VQA) and natural language for visual reasoning (NLVR) |
| Approach: | They propose a cross-modality relevance module that is used in an end-to-end framework to learn the relevance representation between components of various input modalities under supervision of a target task. |
| Outcome: | The proposed approach shows competitive performance on two different language and vision tasks using public benchmarks and improves the state-of-the-art published results. |